Journal of Medical Imaging
● SPIE-Intl Soc Optical Eng
Preprints posted in the last 30 days, ranked by how well they match Journal of Medical Imaging's content profile, based on 11 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Chau, G. N.; Biswas, B. A.; Wagle, B. R.; Maeder, M. E.; Yu, J. B.; Bhattacharya, I.
Show abstract
Automated lesion segmentation is increasingly central to PSMA PET/CT interpretation, supporting staging, treatment planning, and response assessment at a scale that outpaces available nuclear-medicine expertise. However, automated PSMA-PET/CT whole-body lesion segmentation models are trained on images alone, with no knowledge of where in the body prostate metastases actually tend to occur. Radiologists use clinical domain knowledge of metastatic spread, but its absence in machine learning models produces false positives in anatomically implausible locations and missed lesions in high-risk sites such as the liver. In this work, we explore whether population-level spatial knowledge of metastatic spread can be used to augment deep learning segmentation predictions, and how such a prior should be fused with a network's output, without additional training. We build a data-driven metastasis atlas from 375 expert-annotated whole-body PSMA PET/CT scans and investigate its fusion with a trained segmentation network under a Bayesian framework, in which prediction probabilities from an nnU-Net-based lesion segmentation model serve as the likelihood and the data-driven atlas as the prior. Because metastases occupy only a small fraction of whole-body voxels, the atlas's peak probability is too low, and standard power-scaled or naive Bayesian pooling references lack the tools to deal with this shortcoming. This causes these standard fusion strategies to fail and, in the naive Bayesian case, to sharply degrade performance. We instead derive a calibrated, background-referenced log-odds fusion, one of many possible approaches to combine a population atlas with a deep learning model's predictions, distinct from classical multi-atlas label fusion in that it fuses a single population prior with a trained network's softmax rather than combining several registered atlases. Furthermore, this approach is neutral outside atlas support by construction, reduces exactly to the baseline network when unweighted, and requires no retraining. This atlas fusion significantly improved mean Dice over the baseline nnU-Net on a disjoint internal test set ($+0.011$, Holm-adjusted $p=0.021$) and on an independent external cohort ($+0.0129$, Holm-adjusted $p=3.8\times10^{-16}$), with lesion sensitivity improving from 0.849 to 0.861 internally and Dice improving over baseline in every stratified anatomic region, including the rare, high-risk sites motivating this work, while naive Bayesian pooling degrades performance sharply and power-scaled pooling underperforms it throughout. Our findings suggest that population-level spatial priors can meaningfully augment deep learning predictions in whole-body oncologic segmentation, provided the fusion rule is calibrated to where the prior actually carries signal.
dela Sotta, T.; Saavedra, J. M.; Chang, V.; Xavier, A.; Henriquez, H.; Orellana, Y.; Curimil, J.
Show abstract
Diffusion models achieve high reconstruction quality in low-dose computed tomography (LDCT), but their iterative sampling trajectories impose substantial computational costs. Unlike unconditional generation, paired LDCT reconstruction starts from an image that already contains the anatomy and spatial structure of the standard-dose CT (SDCT) target; reconstruction primarily requires correcting dose-related noise and artifacts. We therefore introduce Residual Endpoint Flow Matching (REFM), an LDCT reconstruction method that learns to transport an LDCT image directly toward its paired SDCT endpoint rather than defining a noise-to-image trajectory. REFM predicts the residual velocity along linear interpolations between both images and supports single-step and multi-step reconstruction using the same trained network. We evaluate five model capacities using 1 to 50 Euler steps against deterministic U-Net and diffusion-based baselines. Across all REFM variants, one-step inference consistently provides the highest reconstruction quality. On the TCIA validation set, REFM Base achieves 50.98 dB PSNR and 0.9865 SSIM at 94.54 fps, compared with 50.92 dB, 0.9847, and 9.26 fps for DDPM-10. REFM Small retains 50.71 dB while increasing throughput to 198.56 fps. Without fine-tuning, REFM Base also matches the 25-step DDPM baseline on the external Mayo Clinic dataset, although DDPM remains stronger on synthetically degraded CRLM images. Thus, our results show that exploiting paired anatomical correspondence enables diffusion-level LDCT reconstruction with a single step reconstruction.
Nielsen, M.; Castelo, A.; Altaie, M.; Bennett, J.; Anthony, A.; Siddiqi, N. S.; Gupta, A. C.; Brock, K. K.; Woodland, M.
Show abstract
Reliable clinical deployment of automated liver segmentation requires mechanisms for detecting failures in rare and previously unseen scenarios. Achieving this goal requires an appropriately calibrated threshold that converts an out-of-distribution (OOD) score into a failure prediction. However, threshold calibration typically relies on expert-labeled failures, creating a substantial annotation burden when failures are rare. Building upon our prior work, which uses Pairwise Surface DSC scores as indicators of segmentation quality, we propose a label-free framework for calibrating OOD score thresholds. First, we fitted a log-t distribution to Pairwise Surface DSC scores from a validation set of 400 internal scans to approximate an in-distribution score distribution. New segmentations were assigned significance scores based on their extremity under this fitted distribution and categorized into Low, Medium, and High Risk review groups using statistically principled cutoffs of 0.25 and 0.05. The fitted log-t distribution provided a strong fit to the observed scores and remained robust to moderate contamination by OOD cases. On an independent test set of 500 internal and external scans, the combined Medium and High Risk categories achieved 100% sensitivity and 79% specificity, whereas the High Risk category alone achieved 78% sensitivity and 96% specificity. These results indicate that clinically meaningful failure detection can be derived from unlabeled data. Our code is available at https://github.com/marshalln7/Label_Free_OOD_Threshold_Selection.
BAI, T.-C.; YEH, S.-C.
Show abstract
Foundation models for chest X-ray interpretation make it possible to adapt specialised visual representations with relatively small trainable modules. We report a retrospective study of Low-Rank Adaptation (LoRA) of Rad-DINO Vision Transformer Base with 14x14 patches (ViT-B/14) for 14-class multi-label classification on the National Institutes of Health (NIH) ChestX-ray14 dataset. The official test labels were accessed during earlier model development and configuration comparisons; consequently, every official-test result in this manuscript is explicitly descriptive and non-confirmatory. We used a patient-disjoint 90/10 split of the official trainval pool (77,988 training and 8,536 validation images) and retained the released 25,596-image test partition. The historically selected all-linear LoRA configuration with safe augmentation and g=37 produced a descriptive test macro AUROC of 0.8462 versus the frozen baseline of 0.8295. Comparisons of target modules, patch-token grids, and a Rad-DINO-specific local query head are reported as retrospective comparisons rather than unbiased model-selection evidence. A confident-learning diagnostic flagged 17,653 of 86,524 trainval images (20.4%); this is a model-based flag rate, not a ground-truth label-error rate. A separate counterfactual relabeling sensitivity analysis, which uses the same model to identify and rescore disagreements, changed the descriptive AUROC to approximately 0.9445 after 6,509 policy-defined flips. This value is not achieved model performance and is not a radiologist-audited label-quality ceiling. We provide a validation-only threshold and artifact protocol for future locked evaluation, but a genuinely untouched holdout and new locked selection are required for a confirmatory headline. The existing Zenodo record contains the 25 publication figures only.
Oyarzun Silva, R.; Hernandez Hernandez, P.
Show abstract
Background. Accurate delineation of the gross tumour volume (GTV) - primary tumour (GTVp) and nodal disease (GTVn) - on FDG-PET/CT is a critical step of head and neck radiotherapy planning. Comparisons between lightweight custom networks and the auto-configured nnU-Net v2 are usually reported as end-to-end pipelines, conflating the contribution of the network with that of the inference-time post-processing applied on top of it. We separated the two. Methods. MiniUNet3D (custom 3D U-Net, 18.3 M parameters) and nnU-Net v2 (3d_fullres, 88.2 M parameters) were trained on the same 578 FDG-PET/CT cases (85/15 author-defined split of the HECKTOR 2025 Task 1 set, 8 centres) and evaluated on the same internal cohort. Three arms were compared pairwise: MiniUNet3D raw output at a fixed 0.5 threshold, MiniUNet3D with a locked adaptive post-processing pipeline, and nnU-Net v2. Comparisons used paired Wilcoxon tests with bootstrap confidence intervals, Bonferroni and Benjamini-Hochberg correction, and Cohen's d; catastrophic failure (Dice < 0.01) was compared with an exact McNemar test. Cases with an empty reference for a given target were excluded from that target's analysis (n = 98 GTVp, n = 93 GTVn). Results. With post-processing matched off, nnU-Net v2 was superior: median GTVp Dice 0.799 versus 0.592 (mean difference -0.244, 95 % CI -0.300 to -0.191; d = -0.88) and GTVn 0.774 versus 0.598 (d = -0.82). Post-processing raised MiniUNet3D to 0.800 (GTVp) and 0.738 (GTVn), recovering 79 % of that difference. Post-processed, MiniUNet3D matched nnU-Net v2 on GTVp Dice (p = 0.113) but remained inferior on nodal disease after Bonferroni correction (Dice p = 0.041; surface Dice p = 0.049). Catastrophic GTVp failures were 25/98 raw, 8/98 post-processed and 1/98 for nnU-Net v2 (McNemar p = 0.016). Inference took 34 s versus 78 s per case on the same GPU. Conclusions. Post-processing recovered most, but not all, of the difference between the two models, and it did not confer robustness: an eight-fold higher rate of empty contours on small primaries persisted, which is the more consequential difference for planning safety. Pipeline comparisons reported without a post-processing ablation risk attributing to a network what post-processing supplied.
Suzuki, M.
Show abstract
Background. Extracellular volume fraction (ECV) derived from contrast-enhanced CT is a validated marker of hepatic fibrosis and has been reported to differ between hepatocellular carcinoma (HCC) and intrahepatic cholangiocarcinoma. In published work it is obtained from a small number of hand-placed two-dimensional regions of interest, and the software that computes it is either tied to one manufacturer's workstation or based on spectral or dual-energy acquisition. We are not aware of an accessible tool that produces voxelwise liver ECV maps from conventional single-energy multiphase CT. Methods. We developed CT ECV Mapper, a scripted 3D Slicer extension with a three-layer architecture whose numerical core imports neither slicer nor vtk and is unit-tested outside 3D Slicer. The interactive application provides two-stage registration that the operator inspects and accepts before any ECV is computed, operator-placed three-dimensional regions of interest, user-adjustable calculation parameters, a voxelwise ECV color map and ROI statistics; the same logic layer can be driven unattended across a cohort. The tool was applied to the 164 patients of the public WAW-TACE multiphase HCC/TACE dataset that have both unenhanced and delayed-phase series. Results. 156 of 164 cases (95.1%) completed unattended. Whole-liver ECV had a median of 36.2% (interquartile range 31.9-41.5), consistent with published CT-ECV values for fibrotic and cirrhotic liver. Registering the arterial and portal phases on demand extended tumor ECV from the 38 lesions a conventional two-phase pipeline can reach to 248 lesions in 156 patients. Every failure was attributable to an identifiable mechanism: craniocaudal field-of-view mismatch between phases in six cases, aortic calcification within the blood-pool region in one, and in one case a labeling error in the source dataset, in which the series declared as unenhanced proved to be a second reconstruction of the portal venous phase; this was detected by the blood-pool validity check rather than by visual review. Conclusions. Voxelwise CT ECV mapping of the liver and of hepatic tumors is feasible from conventional multiphase CT on an open platform, both interactively and as an unattended batch, with quality-control instrumentation that fails explicitly and diagnosably. This is a technical development and feasibility report; the application has not been evaluated against a reference standard and no claim of clinical validity is made.
Sandvold, O. F.; Proksa, R.; Perkins, A. E.; Daerr, H.; Koehler, T.; Jacob, T.; Brown, K. M.; Roessl, E.; Noël, P. B.
Show abstract
Spectral computed tomography (CT) is a burgeoning quantitative imaging technique with applications in oncologic diagnostics, prognostic prediction, tissue perfusion studies, and treatment follow-up. While normalized iodine concentration values have been correlated with microenvironmental biophysical changes, obtaining accurate iodine concentrations, particularly at low concentrations remains difficult due to varying spectral CT instrumentation performance. Hybrid spectral CT systems, combining multiple spectral CT instrumentation techniques, address these quantitation insufficiencies by increasing spectral separation but have not been evaluated on a clinically analogous platform. We validate a hybrid spectral CT system, comprised of clinical-grade components, acquiring four distinct effective spectra and applying efficient noise-reducing weighting schemes to compare iodine noise and bias against conventional kVp-Switching (kVp-S). Two tube current levels (50, 350 mA) and three duty cycle ratios (33/67, 50/50, 75/25) were implemented to elucidate radiation dose exposure and kVp-S parameterization impact. A standard quality assurance (QA) and patient-derived, abdominal IodinePrint phantom were scanned on the system. The average absolute bias in iodine density images of the QA phantom was comparable across acquisition techniques, below 0.5 mg/mL, while quantitative noise improved by 22% using noise-optimized weighting schemes. In the IodinePrint phantom aorta and pancreas structures, the noise-optimized weighting scheme increased signal-to-noise ratio (SNR) by 1.3x compared to kVp-S alone. These results highlight the increased precision of hybrid, multi-channel spectral CT systems and motivate CT designs that enable robust CT biomarker development.
Takeuchi, T.; Nomiya, A.
Show abstract
Background: A 2019 report from our institution described a multilayer artificial neural network (ANN) for predicting prostate cancer at biopsy in 334 patients, trained with TensorFlow 1.x and evaluated at three fixed step counts without separating hyperparameter selection from test evaluation. We re-analyzed an expanded cohort from the same institution using contemporary machine-learning practice. Methods: We pooled all available biopsy episodes from the same institutional database (n = 526; 524 after excluding one non-binary outcome code and one record with missing digital rectal examination [DRE] data), retaining the same seven predictors used in the original report (age, prior biopsy history, PSA, prostate volume, DRE, and MRI diffusion-weighted imaging findings in the peripheral and transition zones). Because 27 patients contributed more than one biopsy episode, we used patient-ID-grouped, stratified k-fold cross-validation (StratifiedGroupKFold; scikit-learn 1.8.0) with 3 and 5 folds, repeated over 10 random partitions, to avoid leakage between folds. Four classifiers were compared: L2-regularized logistic regression, gradient boosting, random forest, and a shallow (single hidden layer) multilayer perceptron. Two outcomes were modeled: detection of any prostate cancer, and detection of clinically significant prostate cancer (Gleason score [≥] 7). Results: Any-cancer prevalence was 55.7% (292/524) and Gleason score [≥] 7 prevalence was 39.7% (208/524). With repeated 5-fold cross-validation, gradient boosting gave the highest discrimination for any prostate cancer (mean AUC 0.826, 95% CI 0.823-0.830) and for Gleason score [≥] 7 (mean AUC 0.855, 95% CI 0.852-0.859), closely followed by random forest and logistic regression (AUC 0.81-0.85). The shallow multilayer perceptron performed worse and less consistently than the other three models (any-cancer AUC 0.671; Gleason score [≥] 7 AUC 0.742) and than the deeper five-hidden-layer ANN reported in 2019. Results with 3-fold cross-validation were essentially unchanged. Conclusions: In an expanded cohort, regularized logistic regression, gradient boosting, and random forest all discriminated prostate cancer at biopsy at least as well as the previously reported multilayer ANN, using far simpler models and a methodology that separates hyperparameter tuning from performance estimation. A shallow neural network offered no advantage over these simpler alternatives in this sample size. This is a preprint; the study has not undergone external peer review.
BAI, T.-C.; YEH, S.-C.
Show abstract
CXR report generation may require a vision-language model (VLM) to produce both textual findings and spatial bounding boxes. Generative 4B-7B VLMs can emit non-empty outputs on normal images and empty outputs on abnormal images, motivating explicit structural routing. To evaluate whether a hard inference-time gate before a probabilistic VLM changes output-presence performance and to identify the mechanisms underlying paired STRUCT outcomes. We evaluated CXRxVLM v2, combining a frozen microsoft/rad-dino ViT-B/14 encoder with a 768[->]1 logistic probe (threshold 0.0557) and google/medgemma-4b-it with the pamessina/medgemma-4b-it-cure LoRA adapter. A seed=42 stratified cohort of 500 VinDr-CXR train-pool images (250 NORMAL, 250 ABNORMAL) was compared with Lingshu-7B A_baseline and D_fewshot configurations. Exact paired McNemar tests and stratum-level output-presence analyses were prespecified for the primary configurations; MedGemma 1.5 SigLIP was exploratory. CURE achieved STRUCT = 78.0% (390/500; Wilson 95% CI 74.2-81.4), versus 73.8% for Lingshu A_baseline and 74.2% for D_fewshot. Pairwise p-values were 0.0778, 0.1042, and 0.8642. The paired decomposition showed CURE ABNORMAL non-empty-output advantage of +13.6 percentage points versus Lingshu A (p = 0.0012; +14.0 points versus D, p = 0.0007), while Lingshu had higher NORMAL empty-output rates (+5.2 to +6.4 points; p = 0.0106 and p = 0.0004). The full pipeline used 8.87 GB VRAM and 4.92 s/image mean latency; 53% of records used a 25.7 ms warm gate-negative path after model loading. Equivalent overall STRUCT scores concealed two mechanistically different output regimes: CURE favored ABNORMAL non-empty outputs, whereas Lingshu favored NORMAL empty outputs. This paired decomposition, rather than the aggregate score alone, characterizes how hard-gated and probabilistic systems route output presence.
Do, H. P.; Bekku, M.; Berkeley, D.; Golden, M.; Kitane, S.; Uike, M.; Shinoda, K.; Takayanagi, R.; Takai, H.; Kawai, T.; Seballos, K.; Conley, R.; Sorfleet, K.; Devries, D.; Tymkiw, B.; AlGhuraibawi, W.; Caruthers, S. D.; Kadbi, M.; Provencher, M.; Tashman, S.; Ho, C. P.
Show abstract
Purpose: To determine the feasibility of a 2-minute multi-echo UTE (mecho-UTE) for CT-like bone-weighted contrast and T2* quantification of tissues with short T2/T2*. Methods: Mecho-UTE data acquired from four patients and five healthy subjects were used to assess image quality of the CT-like contrast. All data were reconstructed using conventional gridding (GRID+CONV) and compared with those reconstructed using conjugate gradient SENSE combined with deep learning-based denoising (CG+DLR). Image resolution and sharpness of the CT-like images were assessed using the full width at half maximum (FWHM) and relative edge sharpness (RESH), respectively. Calimetrix UTE-T2* phantom was used to assess the accuracy of T2* quantification of the mecho-UTE sequence. Results: Two-minute mecho-UTE with CG+DLR has similar accuracy (0.37 {+/-} 0.27 vs. 0.67 {+/-} 0.54 ms, p=0.20) and better precision (0.28 {+/-} 0.16 vs. 1.23 {+/-} 0.29 ms, p<0.001) compared to the 5-minute mecho-UTE with GRID+CONV. The 2-minute mecho-UTE with CG+DLR has higher resolution and sharpness compared to the 5-minute scan with GRID+CONV. Conclusion: It is feasible to achieve simultaneous CT-like contrast and T2* quantification of short-T2 tissues in two minutes. When appropriately used, it may simplify logistics, reduce costs, and eliminate radiation exposure risks.
Bourne, R. M.; Arhatari, B.; Watson, G.; Gureyev, T.; Phipps, A.; Dowland, S.; Kurniawan, N.; Sved, P.
Show abstract
Formalin-fixed prostate tissue samples were imaged by propagation-based synchrotron phase contrast micro computed tomography ({micro}CT) with a 3D spatial resolution of ca. 3 {micro}m. Post-{micro}CT, samples were prepared for histology with sections close to coplanar with the transverse {micro}CT image planes. Haematoxylin and eosin stained sections were examined by an expert prostate histopathologist and compared qualitatively with corresponding {micro}CT-visible microstructure features. There is potential for {micro}CT to provide complimentary information to conventional histology and light microscopy without the need for preparation of stained thin sections. For the imaging conditions and spatial resolution of our study, {micro}CT may provide tissue architectural features similar to those used in Gleason grading, albeit without clear subcellular microstructure detail. At the spatial resolution of our study {micro}CT may provide novel 3D microstructure information for validation of diffusion weighted magnetic resonance imaging (MRI) methods. As an example, we demonstrate a qualitative correlation between {micro}CT-derived stromal fibre orientation and preferential water diffusion direction measured by diffusion tensor MRI microscopy of the same sample.
Zareian, B.; Fontaine, K.; Bini, J.
Show abstract
Background. Roughly, half of new type 1 diabetes (T1D) diagnoses occur in individuals under 18 years old and represent a more aggressive destruction of beta cell mass (BCM). [11C]-(+)-PHNO positron emission tomography (PET) imaging is used to assess BCM, but current pancreas PET imaging protocols are limited to adults. Previously published full count data from six healthy controls and five T1Ds (6M/5F; 22 to 53 years old) were used for retrospective analysis. Dynamic [11C]-(+)-PHNO PET/CT scans were acquired and reconstructed using full-count list-mode data. For the current comparison to full count data, 50%, 25% and 10% down-sampled count data were re-reconstructed. Pancreas and spleen (reference region) time-activity-curves (TACs) were assessed, and volume of distribution (VT, mL/cm3) was estimated using the reversible 1-tissue compartment model (1TC) with tmax of 30 min for all count levels. Pancreas and Spleen VT estimates (1TC; tmax= 30 min) were used to calculate non-displaceable binding potential (BPND) and were then correlated to semi-quantitative methods of standardized uptake value ratio (SUVR-1) (20-30 min; ref: spleen) to examine simplified methods using simulated low dose protocols. Finally, we performed dosimetry in adult, adolescent and pediatric phantoms to assess radiation dose for simulated low-dose protocols. Results. Qualitatively, increasing noise can be visualized at successive reduced-count levels images, compared to full-count images. Despite progressively increasing noise in reduced-count images, TACs at each reduced-count level remained similar to full-count TACs in both HC and individuals with T1D. Quantitatively, 1TC VT estimates were similar for all reduced count levels and range of tmax values, compared to full-count (all R2[≥]0.99). Pancreas SUVR-1 (20-30 min) and pancreas BPND (tmax = 30; ref: spleen) were highly correlated for all count levels (all R2[≥]0.80). All age groups were under both the yearly occupational and research scan radiation dose limits when examining mean effective dose equivalent with reduced (1/10th) injected dose protocols. Conclusion. Low-count reconstructed data and simplified reference region approaches provide accurate quantification compared to full-count reconstructions. These results provide evidence that it is possible to perform accurate quantification using simulated low dose protocols to quantify BCM for use in individuals with T1D under 18 years old.
Khan, M. H.; Marin-Pardo, O.; Chakraborty, S.; Lee, K.; Lee, S. Y.; Raman, N.; Iglesias, J. E.; Liew, S.-L.
Show abstract
Accurate stroke lesion segmentation is essential for large-scale neuroimaging studies, yet manual delineation remains labor-intensive, and existing automated methods often struggle to generalize across imaging protocols and stages of recovery. We developed MAESTRO, a deep learning framework for automated lesion segmentation across the stroke recovery continuum using T1-weighted (T1) MRI alone. We hypothesized that combining a transformer-based architecture with an image augmentation strategy would improve segmentation accuracy and robustness under heterogeneous imaging conditions. T1 MRI scans and expert-traced lesion masks from 955 stroke participants across 33 international cohorts were used to train and evaluate MAESTRO within the open-source nnU-Net framework. Performance was evaluated on a held-out test set using spatial and volumetric agreement metrics. An exploratory human-in-the-loop (HITL) evaluation compared correction of MAESTRO-generated segmentations with manual tracing from scratch. MAESTRO achieved the strongest performance across several evaluated model configurations, providing the most accurate lesion localization and lesion volume estimates (median Dice = 0.686; Pearson r = 0.861; ICC = 0.792). Segmentation performance was sensitive to lesion size and stroke chronicity but remained robust across diverse imaging conditions. Additionally, using a HITL workflow to correct MAESTRO segmentations reduced annotation time by 47.4% compared to manual tracing while improving accuracy relative to both automated and manual workflows. MAESTRO is publicly available to enable robust, automated stroke lesion segmentation from T1 MRI. When combined with human review and correction, MAESTRO offers a practical approach for generating standardized, high-quality lesion annotations, helping reduce a major practical barrier to large-scale stroke imaging studies.
Reddy Chimmula, R.; Yong, C.; Love, H. L.; Shiradkar, R.; Holmes, J.; Nair, V.; Tann, M.; Bahler, C.; Oderinde, O. M.
Show abstract
Background: Biochemical recurrence (BCR) occurs in up to 40% of men following radical prostatectomy (RP). Current risk models rely primarily on clinicopathologic variables and may not fully capture the biological heterogeneity associated with recurrence. The Decipher Genomic Classifier (DGC), prostate-specific membrane antigen positron emission tomography (PSMA-PET), and multiparametric magnetic resonance imaging (mpMRI) provide complementary prognostic information that may improve prediction. Objective: To develop and evaluate machine learning (ML) models integrating DGC, PSMA-PET, and mpMRI for preoperative prediction of BCR following RP. Methods: This retrospective study included patients with available preoperative DGC, PSMA-PET, mpMRI, and clinicopathologic data. Logistic regression (LR), random forest (RF), and XGBoost models were developed using single- and multimodality feature combinations. Early- and intermediate-fusion strategies were evaluated. Performance was assessed using an area under the receiver operating characteristic curve (AUC) and accuracy. Clinical utility was evaluated using decision curve analysis. Results: XGBoost consistently outperformed LR and RF. DGC achieved the highest single-modality performance (AUC 0.94, accuracy 86.7%). Among multimodal models, DGC combined with PSMA-PET using intermediate fusion achieved the best overall performance (AUC 0.93, accuracy 87.0%). Addition of mpMRI reduced performance (AUC 0.85, accuracy 83.0%). Decision curve analysis demonstrated positive net benefit across clinically relevant thresholds. Conclusion: XGBoost-based multimodal fusion improved preoperative BCR prediction following RP. DGC was the strongest individual predictor, while integration with PSMA-PET provided the best overall performance, supporting the potential of radiogenomic ML models for personalized risk stratification.
Kim, J.; Kim, B.-s.; Ko, J. S.; Dong, J.; Youn, S. Y.; Jang, J.; Ahn, K.-J.
Show abstract
Purpose Open-source vision-language models (VLMs) can be locally deployed without external internet access, potentially enhancing data security. This study compared the diagnostic performance of general-purpose and medical-purpose open-source VLMs and evaluated their ability to characterize brain metastases on contrast-enhanced (CE) MRI. Materials and Methods Sixty lesion-positive axial CE T1-weighted images and sixty matched lesion-negative images from 60 patients were analyzed using three general-purpose VLMs-InternVL3-8B, Qwen2.5-VL-7B-Instruct, and MiniCPM-V-4.5-and three medical-purpose VLMs-MedGemma-4B-it, LLaVA-Med v1.5, and HuatuoGPT-Vision-7B. Lesion detection performance was assessed using sensitivity, specificity, and balanced accuracy. On lesion-positive images, accuracy was evaluated for lesion count, laterality, anatomic location, enhancement pattern, necrosis, vasogenic edema, and mass effect. Model differences were assessed using Cochran's Q tests followed by pairwise McNemar tests with Benjamini-Hochberg correction. Results The median age of the study patients was 67 years (IQR, 61.0-70.5 years), and 35 patients were male (58.3%). MiniCPM-V-4.5 showed the most balanced diagnostic performance, with a sensitivity of 78.3% (95% CI, 66.4-86.9%) and a specificity of 85.0% (95% CI, 73.9-91.9%), and significantly higher balanced accuracy than all other models. Significant overall differences were observed for lesion count, laterality, location, enhancement pattern, necrosis, and mass effect, but not for vasogenic edema (FDR-adjusted P = 0.056). HuatuoGPT-Vision-7B and MedGemma-4B-it showed relatively consistent accuracy across multiple image assessment tasks, although their performance remained modest. Conclusion Our study demonstrated substantial heterogeneity in the performance of open-source VLMs in brain metastasis evaluation, and medical-purpose VLMs did not outperform general-purpose VLMs.
Dack, E.; Dai, C.; Hoppe, H.; Krueselmann, P.; Meiler, S.; Jutidamrongphan, W.; Wang, L.; Tang, K.
Show abstract
AI-assisted diagnostic tools typically act as a "second opinion," providing radiologists with a discrete prediction or probability score that can be consulted alongside clinical context. This treats AI as an independent advisor rather than a collaborative partner, leaving its reasoning largely opaque. We explore a complementary approach grounded in human-AI collaboration through visual interpretability. Specifically, we investigate (1) radiologist performance when diagnosing chest X-rays from images alone, and (2) whether deep learning-generated heatmaps can support radiologists during this diagnostic process, rather than merely validating a final answer. We developed an interactive application that enables readers to engage directly with model-generated heatmaps as they form their diagnoses, and conducted a user study to evaluate how this influences diagnostic behaviour and accuracy. Our findings offer new insights into integrating interpretable, spatially grounded AI feedback into radiologist workflows. Code, datasets, and the application can be found at https://github.com/eedack01/heatmap_assisted_diagnosis.
Ong, J.; Lau, R.; Chow, K. M.; Huned, D.; Teo, R.; Lee, H. J.; Lim, E. J.; Aslim, E.; Lim, Y. W.; Chen, K.; Tan, Y. Q.; Park, J. J.; Tung, J.
Show abstract
Introduction Anatomical endoscopic enucleation of the prostate (AEEP) techniques, including bipolar enucleation (B-TUEP), holmium laser enucleation (HoLEP), thulium laser enucleation (ThuLEP), and thulium fibre laser enucleation (ThuFLEP), demonstrate comparable clinical outcomes for benign prostatic hyperplasia. As clinical equivalence is increasingly established, cost becomes a key determinant of modality selection. We performed a cost minimisation analysis comparing index procedural costs across AEEP modalities from an institutional perspective. Methods A cost minimisation model was developed from the institutional perspective, incorporating amortised capital costs, maintenance, and consumables. In addition to the base-case scenario of 180 cases per year, we modelled two additional case volume scenarios: low (50 cases/year) and high (500 cases/year) volume. Thu:YAG laser fibres were modelled on two scenarios: disposable single-use, and reusable fibres (up to 10 cases per fibre). Breakeven analysis determined the threshold volume at which each laser modality achieves cost parity with B-TUEP, and one-way sensitivity analysis was performed on key cost parameters. Analysis was limited to index procedural costs calculated in Singapore dollars. Results At the base case of 180 cases per year, B-TUEP had the lowest index procedure cost (SGD 1,018), followed by ThuFLEP (SGD 1,584), ThuLEP (1,599), and HoLEP (SGD 1,655). Breakeven analysis demonstrated that HoLEP, ThuLEP, and ThuFLEP can never achieve cost parity with B-TUEP when laser fibres are single-use, as laser modalities carry higher costs on both capital and per-case dimensions. ThuLEP with reusable fibres (10 uses per fibre) was the only modality to cross below B-TUEP, at a breakeven volume of 198 cases per year. At 500 cases per year with reusable fibres, ThuLEP achieved the lowest cost (SGD 847), representing a 15.4% saving over B-TUEP. Sensitivity analysis identified annual case volume and B-TUEP loop cost as the most influential parameters. Conclusion Index procedural costs in AEEP are strongly influenced by case volume and consumable strategy. While B-TUEP remains cost-efficient at low volume, high-volume practice combined with reusable Thu:YAG fibre technology enables cost parity and potential cost advantage for laser enucleation. These findings highlight the importance of economies of scale and device utilisation in technology adoption.
Calado, A.; de Almeida, J. G.; Verde, A. S. C.; Tsiknakis, M.; Marias, K.; Regge, D.; Papanikolaou, N.; ProCAncer-I Consortium,
Show abstract
Purpose: To prospectively validate a semi-supervised learning framework with a lesion-only teacher model (RG-SSL-LOC) for scalable clinically significant prostate cancer detection on biparametric MRI (bpMRI) and assess its added value in multimodal models. Materials and Methods: A multicenter dataset of 13,706 bpMRI examinations (13,630 patients, 27 centers) was used for model development/validation. Three segmentation models (fully supervised learning [FSL], a state-of-the-art report-guided semi-supervised approach [RG-SSL], and the proposed RG-SSL-LOC) were evaluated at lesion- and case-level on external retrospective, external prospective, and internal prospective cohorts. Predictions from the best-performing model were combined with clinico-radiologic variables in a multimodal approach. All case-level results were compared with PI-RADS. Results: At lesion level, RG-SSL-LOC achieved higher median Dice than FSL and RG-SSL (0.49 vs 0.41 and 0.40; both p<.001). At case level, RG-SSL-LOC achieved area-under-the-curve (AUC) values of 0.83, 0.82, and 0.87 in the external retrospective, external prospective, and internal prospective cohorts, respectively. Compared with FSL, AUCs were 0.84 (p=.237), 0.80 (p=.020), and 0.84 (p<.001); compared with RG-SSL, AUCs were 0.83 (p=.929), 0.82 (p=.652), and 0.86 (p=.007); compared with PI-RADS, AUCs were 0.78 (p=.055), 0.83 (p=.652) and 0.86 (p=.480). Combined with clinico-radiological variables, RG-SSL-LOC significantly improved AUC versus clinico-radiological variables alone in the external retrospective (0.85 vs 0.80, p=.002), external prospective (0.87 vs 0.84, p=.008), and internal prospective (0.91 vs 0.88, p<.001) cohorts; in the latter, it reduced unnecessary biopsies by 15.19%. Conclusion: RG-SSL-LOC achieves better segmentation quality than other methods, demonstrates robust prospective multicenter performance and improves multimodal detection.
Ye, Z.; He, F.; Zhao, T.; Xia, W.
Show abstract
Ultrathin endoscopy is highly attractive for real-time tissue imaging in narrow and hard-to-reach regions of the body. A single multimode fibre (MMF) is an attractive probe because of its small diameter, flexibility, and diffraction-limited spatial resolution enabled by the large number of transverse modes guided within a single core. Because the distal fibre tip is inaccessible during endoscopy, reflection-mode imaging, in which the same fibre delivers illumination and collects backscattered light, is more practical than transmission-mode imaging. However, image recovery from the resulting speckle pattern is challenging because light undergoes double-pass propagation through the MMF, with mode coupling and dispersion; the backscattered signal is weak, and the camera records intensity only, without phase information. Here, we propose a single-shot reflection-mode MMF imaging framework that combines a reflected real-valued intensity transmission matrix (reflected-RVITM) with an image restoration network. The reflected-RVITM is calibrated using intensity-only measurements, without interferometry or phase retrieval, and provides a physics-guided initial reconstruction from a single backscattered speckle frame. A restoration network then refines this initial reconstruction instead of inverting the raw speckle. Four restoration backbones are evaluated: HPM-Attention-UNet, GAM, MambaIRv2, and CICPNet. On matched datasets, hybrid models outperformed corresponding networks trained to map raw speckle directly to images. For example, HPM-Attention-UNet on MNIST improved mean PCC from 0.572 to 0.944 (+65.1%). Under domain shift, with training only on Fashion-MNIST and tested on unseen CIFAR scenes, hybrid models achieved mean PCC of 0.61-0.65, compared with 0.36-0.50 for direct learning. This framework is further demonstrated using physical objects at the distal fibre tip. These results demonstrate that a reflected-RVITM physics prior combined with a restoration network enables single-shot image recovery after intensity-only calibration, offering a phase-retrieval-free and generalisable route towards minimally invasive reflection-mode MMF endoscopy.
Mukherjee, S.; Templeton, K. A.; Schiff, S. J.; Monga, V.
Show abstract
Objective: Accurate volumetric analysis of the brain and cerebrospinal fluid (CSF) is essential for monitoring hydrocephalus, a significant pediatric neurological condition. While computed tomography (CT) provides high-quality volumetric assessment, its associated ionizing radiation poses risks, especially for children. Low-field magnetic resonance imaging (LF-MRI) offers a safer and more accessible alternative, particularly in resource-constrained settings. However, its lower resolution and increased susceptibility to structural distortions make accurate segmentation challenging. This study aims to demonstrate that reliable volumetric measurements can be obtained from LF-MRI that are comparable to CT, enabling safer and more frequent monitoring of infants with hydrocephalus. Approach: We propose EnSegNet-Cross, a cross-modality, enhancement-aware segmentation network for brain volume analysis using LF-MRI. The framework leverages high-fidelity CT data during training but requires only LF-MRI during inference. At the core of the framework is a novel cross-modal topological penalty designed to minimize discrepancies between predicted LF-MRI and CT structures. A central contribution is the integration of a three-dimensional topological loss based on persistent homology, which penalizes topological discrepancies in CSF regions, specifically CSF holes formed by enclosed brain parenchyma, between CT and LF-MRI segmentations. Incorporating these structural priors facilitates generalization across heterogeneous clinical cases while eliminating the need for CT data during inference, resulting in more anatomically coherent and topologically faithful segmentations. Main Results: On a curated cohort of infants with hydrocephalus who had paired LF-MRI and CT scans, including cases with infectious and non-infectious causes, EnSegNet-Cross consistently outperformed state-of-the-art machine learning alternatives. It achieved the highest Dice score of 0.8532 plus/minus 0.03 and Volume Score of 0.9318 plus/minus 0.03. The method also demonstrated robust performance in challenging cases with confounding factors, achieving a Dice score of 0.8340 plus/minus 0.03 and a Volume Score of 0.9111 plus/minus 0.05. By leveraging CT-derived topological priors, EnSegNet-Cross successfully handled anatomically complex scenarios in which conventional models failed. Significance: EnSegNet-Cross provides a reliable and interpretable solution for brain and CSF segmentation, particularly in complex cases of hydrocephalus. This study demonstrates that high-fidelity volumetric estimates can be achieved using only LF-MRI, facilitating frequent, radiation-free monitoring. By bridging the fidelity gap between low-quality LF-MRI and high-resolution CT through clinically grounded enhancement and topological supervision, EnSegNet-Cross offers a robust clinical tool for brain volumetric analysis in infants with hydrocephalus using LF-MRI.